Serveur d'exploration sur l'OCR

Attention, ce site est en cours de développement !
Attention, site généré par des moyens informatiques à partir de corpus bruts.
Les informations ne sont donc pas validées.

A conclusive methodology for rating OCR performance

Identifieur interne : 001348 ( Main/Exploration ); précédent : 001347; suivant : 001349

A conclusive methodology for rating OCR performance

Auteurs : Nathan E. Brener [États-Unis] ; S. S. Iyengar [États-Unis] ; O. S. Pianykh [États-Unis]

Source :

RBID : ISTEX:8B4039C936354C5EA2A1027A0E8248A4629FAC7B

Descripteurs français

English descriptors

Abstract

One of the most challenging topics in the automatic document rating process is the development of a rating scheme for the image quality of documents. As part of the Department of Energy (DOE) document declassification program, we have developed a generalized rating system to predict the optical character recognition (OCR) accuracy level that is achieved when processing a document. The need for such a system emerged from the declassification of degraded, typewriter‐era documents, which is currently a time‐consuming manual process. This article presents the statistical analysis of the most influential document quality features affecting OCR accuracy, develops consistent predictive models for four currently used OCR engines, and studies the applicability of different OCR products to the DOE document declassification process. This study is expected to lead to an efficient and completely automated document declassification system.

Url:
DOI: 10.1002/asi.20214


Affiliations:


Links toward previous steps (curation, corpus...)


Le document en format XML

<record>
<TEI wicri:istexFullTextTei="biblStruct">
<teiHeader>
<fileDesc>
<titleStmt>
<title xml:lang="en">A conclusive methodology for rating OCR performance</title>
<author>
<name sortKey="Brener, Nathan E" sort="Brener, Nathan E" uniqKey="Brener N" first="Nathan E." last="Brener">Nathan E. Brener</name>
</author>
<author>
<name sortKey="Iyengar, S S" sort="Iyengar, S S" uniqKey="Iyengar S" first="S. S." last="Iyengar">S. S. Iyengar</name>
</author>
<author>
<name sortKey="Pianykh, O S" sort="Pianykh, O S" uniqKey="Pianykh O" first="O. S." last="Pianykh">O. S. Pianykh</name>
</author>
</titleStmt>
<publicationStmt>
<idno type="wicri:source">ISTEX</idno>
<idno type="RBID">ISTEX:8B4039C936354C5EA2A1027A0E8248A4629FAC7B</idno>
<date when="2005" year="2005">2005</date>
<idno type="doi">10.1002/asi.20214</idno>
<idno type="url">https://api.istex.fr/document/8B4039C936354C5EA2A1027A0E8248A4629FAC7B/fulltext/pdf</idno>
<idno type="wicri:Area/Istex/Corpus">000000</idno>
<idno type="wicri:Area/Istex/Curation">000000</idno>
<idno type="wicri:Area/Istex/Checkpoint">000C57</idno>
<idno type="wicri:doubleKey">1532-2882:2005:Brener N:a:conclusive:methodology</idno>
<idno type="wicri:Area/Main/Merge">001384</idno>
<idno type="wicri:source">INIST</idno>
<idno type="RBID">Pascal:06-0252013</idno>
<idno type="wicri:Area/PascalFrancis/Corpus">000391</idno>
<idno type="wicri:Area/PascalFrancis/Curation">000395</idno>
<idno type="wicri:Area/PascalFrancis/Checkpoint">000444</idno>
<idno type="wicri:doubleKey">1532-2882:2005:Brener N:a:conclusive:methodology</idno>
<idno type="wicri:Area/Main/Merge">001474</idno>
<idno type="wicri:Area/Main/Curation">001348</idno>
<idno type="wicri:Area/Main/Exploration">001348</idno>
</publicationStmt>
<sourceDesc>
<biblStruct>
<analytic>
<title level="a" type="main" xml:lang="en">A conclusive methodology for rating OCR performance</title>
<author>
<name sortKey="Brener, Nathan E" sort="Brener, Nathan E" uniqKey="Brener N" first="Nathan E." last="Brener">Nathan E. Brener</name>
<affiliation wicri:level="2">
<country xml:lang="fr">États-Unis</country>
<placeName>
<region type="state">Louisiane</region>
</placeName>
<wicri:cityArea>Department of Computer Science, Louisiana State University, Baton Rouge</wicri:cityArea>
</affiliation>
<affiliation wicri:level="1">
<country wicri:rule="url">États-Unis</country>
</affiliation>
</author>
<author>
<name sortKey="Iyengar, S S" sort="Iyengar, S S" uniqKey="Iyengar S" first="S. S." last="Iyengar">S. S. Iyengar</name>
<affiliation wicri:level="2">
<country xml:lang="fr">États-Unis</country>
<placeName>
<region type="state">Louisiane</region>
</placeName>
<wicri:cityArea>Department of Computer Science, Louisiana State University, Baton Rouge</wicri:cityArea>
</affiliation>
</author>
<author>
<name sortKey="Pianykh, O S" sort="Pianykh, O S" uniqKey="Pianykh O" first="O. S." last="Pianykh">O. S. Pianykh</name>
<affiliation wicri:level="2">
<country xml:lang="fr">États-Unis</country>
<placeName>
<region type="state">Louisiane</region>
</placeName>
<wicri:cityArea>Department of Computer Science, Louisiana State University, Baton Rouge</wicri:cityArea>
</affiliation>
</author>
</analytic>
<monogr></monogr>
<series>
<title level="j">Journal of the American Society for Information Science and Technology</title>
<title level="j" type="abbrev">J. Am. Soc. Inf. Sci.</title>
<idno type="ISSN">1532-2882</idno>
<idno type="eISSN">1532-2890</idno>
<imprint>
<publisher>Wiley Subscription Services, Inc., A Wiley Company</publisher>
<pubPlace>Hoboken</pubPlace>
<date type="published" when="2005-10">2005-10</date>
<biblScope unit="volume">56</biblScope>
<biblScope unit="issue">12</biblScope>
<biblScope unit="page" from="1274">1274</biblScope>
<biblScope unit="page" to="1287">1287</biblScope>
</imprint>
<idno type="ISSN">1532-2882</idno>
</series>
<idno type="istex">8B4039C936354C5EA2A1027A0E8248A4629FAC7B</idno>
<idno type="DOI">10.1002/asi.20214</idno>
<idno type="ArticleID">ASI20214</idno>
</biblStruct>
</sourceDesc>
<seriesStmt>
<idno type="ISSN">1532-2882</idno>
</seriesStmt>
</fileDesc>
<profileDesc>
<textClass>
<keywords scheme="KwdEn" xml:lang="en">
<term>Automatic processing</term>
<term>Document processing</term>
<term>Image evaluation</term>
<term>Image processing</term>
<term>Image quality</term>
<term>Optical character recognition</term>
<term>Parameter estimation</term>
<term>Regression analysis</term>
<term>System performance</term>
</keywords>
<keywords scheme="Pascal" xml:lang="fr">
<term>Analyse régression</term>
<term>Estimation paramètre</term>
<term>Evaluation image</term>
<term>Performance système</term>
<term>Qualité image</term>
<term>Reconnaissance optique caractère</term>
<term>Traitement automatique</term>
<term>Traitement document</term>
<term>Traitement image</term>
</keywords>
</textClass>
<langUsage>
<language ident="en">en</language>
</langUsage>
</profileDesc>
</teiHeader>
<front>
<div type="abstract" xml:lang="en">One of the most challenging topics in the automatic document rating process is the development of a rating scheme for the image quality of documents. As part of the Department of Energy (DOE) document declassification program, we have developed a generalized rating system to predict the optical character recognition (OCR) accuracy level that is achieved when processing a document. The need for such a system emerged from the declassification of degraded, typewriter‐era documents, which is currently a time‐consuming manual process. This article presents the statistical analysis of the most influential document quality features affecting OCR accuracy, develops consistent predictive models for four currently used OCR engines, and studies the applicability of different OCR products to the DOE document declassification process. This study is expected to lead to an efficient and completely automated document declassification system.</div>
</front>
</TEI>
<affiliations>
<list>
<country>
<li>États-Unis</li>
</country>
<region>
<li>Louisiane</li>
</region>
</list>
<tree>
<country name="États-Unis">
<region name="Louisiane">
<name sortKey="Brener, Nathan E" sort="Brener, Nathan E" uniqKey="Brener N" first="Nathan E." last="Brener">Nathan E. Brener</name>
</region>
<name sortKey="Brener, Nathan E" sort="Brener, Nathan E" uniqKey="Brener N" first="Nathan E." last="Brener">Nathan E. Brener</name>
<name sortKey="Iyengar, S S" sort="Iyengar, S S" uniqKey="Iyengar S" first="S. S." last="Iyengar">S. S. Iyengar</name>
<name sortKey="Pianykh, O S" sort="Pianykh, O S" uniqKey="Pianykh O" first="O. S." last="Pianykh">O. S. Pianykh</name>
</country>
</tree>
</affiliations>
</record>

Pour manipuler ce document sous Unix (Dilib)

EXPLOR_STEP=$WICRI_ROOT/Ticri/CIDE/explor/OcrV1/Data/Main/Exploration
HfdSelect -h $EXPLOR_STEP/biblio.hfd -nk 001348 | SxmlIndent | more

Ou

HfdSelect -h $EXPLOR_AREA/Data/Main/Exploration/biblio.hfd -nk 001348 | SxmlIndent | more

Pour mettre un lien sur cette page dans le réseau Wicri

{{Explor lien
   |wiki=    Ticri/CIDE
   |area=    OcrV1
   |flux=    Main
   |étape=   Exploration
   |type=    RBID
   |clé=     ISTEX:8B4039C936354C5EA2A1027A0E8248A4629FAC7B
   |texte=   A conclusive methodology for rating OCR performance
}}

Wicri

This area was generated with Dilib version V0.6.32.
Data generation: Sat Nov 11 16:53:45 2017. Site generation: Mon Mar 11 23:15:16 2024